Aggregating Probabilistic XML

نویسندگان

  • Serge Abiteboul
  • T.-H. Hubert Chan
  • Evgeny Kharlamov
  • Werner Nutt
  • Pierre Senellart
چکیده

Les sources d’incertitude et d’imprécision des données sont nombreuses. Une manière de gérer cette incertitude est d’associer aux données des annotations probabilistes. De nombreux modèles de bases de données probabilistes ont ainsi été proposés, dans les cadres relationnel et semi-structuré. Ce dernier est particulièrement adapté à la gestion de données incertaines provenant de traitement automatiques. Un important problème, dans le cadre des bases de données probabilistes XML, est celui des requêtes d’agrégation (count, sum, avg, etc.), qui n’a pas été étudié jusqu’à présent. Dans un modèle unifiant les différents modèles probabilistes semi-structurés étudiés à ce jour, nous présentons des algorithmes pour calculer la distribution des résultats de l’agrégation (qui exploitent certaines propriétés de régularité des fonctions d’agrégation), ainsi que des moments (en particulier, espérance et variance) de celle-ci. Nous prouvons également l’intractabilité de certains de ces problèmes.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

A Probabilistic Approach to XML Data Management

Uncertainty is ubiquitous in data and can take various forms. Usually, this is not formally taken into account: only the most likely data interpretation is kept for future processing, or all probable choices of correct information above a threshold are maintained. We claim this is not sufficient. There is a need for managing the imprecision in data more rigorously, and the current thesis addres...

متن کامل

An Efficient and Versatile Query Engine for TopX Search

This paper presents a novel engine, coined TopX, for efficient ranked retrieval of XML documents over semistructured but nonschematic data collections. The algorithm follows the paradigm of threshold algorithms for top-k query processing with a focus on inexpensive sequential accesses to index lists and only a few judiciously scheduled random accesses. The difficulties in applying the existing ...

متن کامل

Research on Basic Operations for Query Probabilistic XML Document Based on Path Set

Probabilistic XML tree is a typical probabilistic XML data model for representing uncertainty in the form of probability which can be described the nested probabilistic XML element unit. The path set in probabilistic XML unit tree can be analyzed through parsing it and drawing out the corresponding probabilistic XML unit schema tree. The node probability and the node probability threshold compu...

متن کامل

Research on Querying Node Probability Method in Probabilistic XML Data Based on Possible World

In order to solve the low efficiency problem of directly querying single node probability in the set of all ordinary XML data obtained by enumerating possible world set of the corresponding probabilistic XML data, the method is presented that probabilistic XML data of possible world set is represented by semis-structured information unit. And it is modeled probabilistic XML data tree. Then the ...

متن کامل

Probabilistic XML functional dependencies based on possible world model

With the increase of uncertain data in many new applications, such as sensor network, data integration, web extraction, etc., uncertainty both in relational databases and XML datasets has attracted more and more research interests in recent years. As functional dependencies (FDs) are critical and necessary to schema design and data rectification in relational databases and XML datasets, it is a...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2009